Reinforcement Learning Startups funded by Y Combinator (YC) 2026

July 2026

Browse 30 of the top Reinforcement Learning startups funded by Y Combinator.

We also have a Startup Directory where you can search through over 5,000 companies.

  • Olam Labs
    Olam Labs
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    Evaluating models for social behavior, safety, and performance through multi-agent simulated games. We use popular games like Catan, Risk, or Poker, and make humans come play them against talking AI opponents, but on the backend we're actually running a multi-agent environment used to evaluate the models for skill and behavior over large sample sizes. You can play in Social Arena and view our preliminary benchmarks today at https://olamlabs.ai/
    reinforcement-learning
    gaming
    data-engineering
    ai
  • Belvedir
    Belvedir
    Y Combinator LogoS2026
    Active • 1 employees • San Francisco
    The everyday person should own their intelligence. Our platform trains custom AI models and memory systems, hosts them privately, and continually improves them as they are used.
    reinforcement-learning
    machine-learning
  • Ooak Data
    Ooak Data
    Y Combinator LogoS2026
    Active • 5 employees • Paris, France
    The promise of AI agents is still unmet in complex, real-world business environments. Our mission is to make it possible for agents to reliably accomplish concrete and useful tasks. That’s why we're building **Alexandria,** the world's largest library of real-world business workflow datasets.
    artificial-intelligence
    reinforcement-learning
  • Fabraix
    Fabraix
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    Fabraix builds state-of-the-art AI red-teaming agents that continuously detect security vulnerabilities in customer-facing AI. Our product, Nyx, has already found vulnerabilities in agents at dozens of Fortune 500 companies. On AgentHarm, the leading benchmark for offensive AI security, Nyx achieved a 78% attack success rate, compared with 67% for GPT-5.6 Sol. Companies usually pay for penetration tests one project at a time. They select which systems to include, give a team a few weeks to find vulnerabilities, and receive a report when the project ends. Testing everything this way is very expensive. Globally, penetration tests cover only 26% of the software attack surface. AI agents can change more often than teams can test them. A new model, prompt, tool, permission, or data source can change what an agent does, even when the application code stays the same. The report describes only the version that was tested. AI is also increasing how much software companies produce and how often it changes. We built Nyx to automate this work. It does three things: 1. Nyx connects through the same interfaces customers use. It tests chat, voice, browser, and coding agents without source code or a special integration. 2. It draws on more than 10,000 jailbreaks that we have collected and classified, the largest such library we know of. 3. It uses each response to decide what to try next and can pursue a promising approach for hundreds of turns. In our tests, these adaptive attacks succeeded 20 times as often as attacks that were replayed unchanged or stopped after a fixed number of turns. Nyx often finds its first vulnerability within minutes or hours rather than days or weeks. Because the work is automated, companies can repeat the test with every change and at much lower cost. We're also the team behind ACE (Adversarial Cost to Exploit), a benchmark that measures AI security in terms of how much it costs attackers to break an AI system; providing a game-theoretic framework to understand how motivated a rational attacker would be in exploiting the system.
    cybersecurity
    ai
    reinforcement-learning
  • Praxis AI
    Praxis AI
    Y Combinator LogoS2026
    Active • 3 employees • San Francisco
    Praxis captures the data physical AI is bottlenecked on: egocentric video, 3D scans, and multimodal capture from inside real industrial and residential environments. Embedded within publicly listed and unicorn-scale conglomerates, we reach 60k workers across 5 continents and 150+ environment types, supplying the real-world training data frontier labs and humanoid companies can't source anywhere else.
    reinforcement-learning
    hard-tech
    robotic-process-automation
  • Markov
    Markov
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    We source high quality tasks and data to train the next generation of computer use AI models.
    reinforcement-learning
    data-labeling
    data-engineering
  • Brumby (Formerly GrazeMate)
    Brumby (Formerly GrazeMate)
    Y Combinator LogoW2026
    Active • 3 employees • Sydney NSW, Australia
    Brumby builds autonomous drones that herd cattle. On command, our drones fly to a paddock, position themselves around the mob, and move them where they need to go. What used to take a full day of helicopters, motorbikes, and horses now runs on a schedule. We work with some of the largest cattle ranches in the world. While the drones are herding, they're also estimating animal weights, measuring grass biomass, monitoring water levels, and flagging sick animals. We're building physical AI that lets a grazier manage thousands of head across millions of acres from their phone.
    agriculture
    reinforcement-learning
    computer-vision
    drones
  • Cortex AI
    Cortex AI
    Y Combinator LogoF2025
    Active • 3 employees • San Francisco
    Cortex AI builds the world’s most diverse and large-scale real-world workplace robot & egocentric dataset — where the physical world becomes the next training and evaluation set for embodied AI. We power frontier labs developing robotics foundation models and general-purpose robots by providing the data they need: 1️⃣ Egocentric Data — real-workplace human video with hand/body pose, depth, and subtask labels. 2️⃣ Robot Data — trajectories collected from manipulators and humanoids in real industry settings. 3️⃣ Human-in-the-Loop Rollouts & Evals — real-world deployments with remote operators who recover robots when they fail, capturing data that feeds back into training and continuously improves models. Additionally, through the Cortex Marketplace, workplaces get paid to host data-collection and evaluation sessions, while labs access the in-the-wild data that truly matters. This draws on Lucas’s previous experience as co-founder of Carousell, a C2C marketplace that scaled to a $1B+ valuation.
    robotics
    reinforcement-learning
    artificial-intelligence
  • hillclimb
    hillclimb
    Y Combinator LogoF2025
    Active • 4 employees • San Francisco
    We work with frontier AI labs to help train their agents to become AI research scientists
    reinforcement-learning
  • Topological
    Topological
    Y Combinator LogoS2025
    Active • 2 employees • San Francisco
    Topological is developing physics-based foundation models for CAD optimization. We help hardware teams iterate at the same speed that software teams do. Our technology is accelerating the engineering workflow with AI and scales design and optimization to identify the ideal designs for complex problems given their physical constraints with enhanced speed and performance. Our first model, UToP-v1, is a SOTA topology optimization model that understands physics, geometry, and manufacturability. It can generate the most efficient design given a problem’s physical requirements. It has <5% compliance error and is 1930x faster than current methods. We're reimagining mechanical engineering and computational design with precision spatial AI.
    ai
    3d-printing
    design-tools
    reinforcement-learning
    robotics
  • Monte
    Monte
    Y Combinator LogoS2025
    Active • 3 employees • San Francisco
    Monte is an applied research lab building products for continual learning and recursive self-improvement for AI agents.
    b2b
    infrastructure
    reinforcement-learning
    artificial-intelligence
  • Idler
    Idler
    Y Combinator LogoS2025
    Active • 13 employees • San Francisco
    Idler builds reinforcement learning environments that teach AI models to code at expert human levels. We create training environments based on real-world coding scenarios that prepare models for the complex challenges they'll face in production.
    reinforcement-learning
  • Kairos
    Kairos
    Y Combinator LogoP2025
    Active • 2 employees • San Francisco
    Kairos closes the last-mile reliability gap in AI deployments. We bring frontier techniques to companies in critical industries, deploying specialized agents that encode operator expertise and reliably automate their most manual workflows.
    reinforcement-learning
    aiops
    machine-learning
    ai
  • Aviro
    Aviro
    Y Combinator LogoP2025
    Active • 2 employees • San Francisco
    Aviro is a research partner with four frontier AI labs, F500 companies, and some of the top RL data vendors building post-training datasets. We focus on tasks involving thousands of tool steps for coding, computer use, and knowledge work.
    reinforcement-learning
  • Cartpole
    Cartpole
    Y Combinator LogoP2025
    Active • 2 employees • San Francisco
    We're creating reinforcement learning environments for training frontier models.
    reinforcement-learning
    ml
    artificial-intelligence
    data-labeling
  • Freesolo
    Freesolo
    Y Combinator LogoP2025
    Active • 4 employees • San Francisco
    Freesolo works with companies to encode user-trajectory knowledge into specialized models that outperform the state of the art at much lower latency and cost.
    reinforcement-learning
    b2b
    ai
  • hud
    hud
    Y Combinator LogoW2025
    Active • 15 employees • San Francisco
    HUD (YC W25) is developing agentic evals and RL environments for Computer Use Agents (CUAs) that browse the web for frontier AI labs. Our CUA Evals framework is the first comprehensive evaluation tool for CUAs. People don't actually know if AI agents are working reliably. To make AI agents work in the real world, we need detailed evals for a huge range of tasks. We're backed by Y Combinator, and work closely with frontier AI labs to provide agent evaluation and training infrastructure at scale.
    ai
    reinforcement-learning
  • Agentin AI
    Agentin AI
    Y Combinator LogoW2025
    Active • 2 employees • San Francisco
    At Agentin AI, we build Agents that move data and take actions across enterprise systems, like Salesforce, NetSuite and SAP. These agents are difficult to build because each enterprise heavily customizes their systems but we solved that by training our Agents to learn and adapt from failures, applying reinforcement learning techniques we developed.
    enterprise
    reinforcement-learning
    ai
  • TrainLoop
    TrainLoop
    Y Combinator LogoW2025
    Active • 6 employees • San Francisco
    TrainLoop makes it effortless for developers to supercharge LLM performance through reinforcement learning.
    developer-tools
    generative-ai
    reinforcement-learning
  • Osmosis
    Osmosis
    Y Combinator LogoW2025
    Active • 6 employees • San Francisco
    Osmosis is a post-training platform that helps companies fine-tune language models using reinforcement learning. We work with fast-growing AI companies to train task/domain-specific models that beat foundation models on performance, cost, and latency. Our platform handles compute orchestration, reward modeling, and training run observability as a CLI-based product usable by developers and agents.
    reinforcement-learning
    machine-learning
    infrastructure
    artificial-intelligence
  • Synth
    Synth
    Y Combinator LogoF2024
    Active • 2 employees • San Francisco
    Choose a coding agent harness, model, and task dataset and optimize context and prompts to get the best performance for long-horizon tasks
    ai
    reinforcement-learning
  • Vibrant Labs
    Vibrant Labs
    Y Combinator LogoW2024
    Active • 6 employees • San Francisco
    We work on methods to autonomously scaling evals/environments for post-training agents.
    generative-ai
    open-source
    developer-tools
    ai
    reinforcement-learning
  • JustAI
    JustAI
    Y Combinator LogoW2024
    Active • 4 employees • San Francisco
    Always-on AI agents for 1-1 personalization at scale
    ai
    marketing
    personalization
    reinforcement-learning
    workflow-automation
  • Velos
    Velos
    Y Combinator LogoW2023
    Active • 3 employees • San Francisco
    Velos helps non-technical operations teams automate complex, manual back-office tasks with AI workers instead of overseas teams. Unlike traditional robotic process automation (RPA) platforms like UiPath, Velos automations use machine learning to reliably handle ambiguities in their tasks, eliminating the need for an army of maintenance engineers and consultants to build and maintain your automations. We're automating the repetitive work people hate to do.
    generative-ai
    reinforcement-learning
    artificial-intelligence
    automation
    robotic-process-automation
  • Atmeto
    Atmeto
    Y Combinator LogoW2023
    Active • 2 employees • Los Angeles
    Founded in 2022, Atmeto was started as a place to develop and apply machine learning to solve the world's biggest problem—climate change. Our current priority is getting the grid to run on 100% clean energy, which is currently limited by battery storage (specifically, the algorithms that control them). We're redefining these algorithms to unlock gigawatts of untapped energy storage capacity, enabling the grid to run on more clean energy from wind and solar.
    climate
    climatetech
    reinforcement-learning
    energy-storage
    energy
  • rct AI
    rct AI
    Y Combinator LogoW2019
    Active • 40 employees • Los Angeles
    rct AI is providing AI solutions to the game industry and building the true Metaverse with AI generated content. By using cutting-edge technologies, especially deep learning and reinforcement learning, rct AI creates a truly dynamic and intelligent user experience both on the consumers’ side and production’s side. The founding team ever built a company, Raventech together and helped make it acquired by Baidu (NASDAQ:BIDU) in 2017.
    gaming
    reinforcement-learning
    metaverse
  • Sepal AI
    Sepal AI
    Y Combinator LogoS2024
    Acquired • 15 employees • San Francisco
    Sepal is a data research company on a mission to advance human knowledge and capabilities through safe AI. We partner with the world’s leading AI labs and enterprises to help their models get better at the tasks people actually want them to do. We’ve built a Cloud-Native Agent Dataset Factory which turns the process of generating evaluation and training data from manual, inconsistent, and labor-intensive into something automated, standardized, and scalable. Sepal AI was founded in 2024 by engineers and operators from Vercel and Turing. We went through Y Combinator, raised several million dollars from leading investors, and already count multiple Fortune 500s and top AI research labs as paying customers.
    data-labeling
    aiops
    reinforcement-learning
    ai
  • Resonance
    Resonance
    Y Combinator LogoW2024
    Acquired • 2 employees • San Francisco
    Resonance hyper-personalizes MarTech campaign content and automatically refreshes and stores high performing content for re-use.
    reinforcement-learning
    artificial-intelligence
    saas
    subscriptions
    marketing